feat(compile): torch.compile support for flagos as a first-class inductor GPU device - #41
Merged
Merged
Conversation
Wires torch.compile into the flagos (PrivateUse1) backend by registering flagos with TorchInductor as a real GPU device, so the traced graph is handed to compile_fx unchanged and inductor emits Triton kernels that operate on flagos tensors directly. The earlier approach rewrote the graph and its example inputs to cuda before compiling. That is not just a copy per call: at::getAccelerator() is PrivateUse1/flagos in this build, and torch::autograd::Node::stream() only yields a stream when a node's input device type equals the accelerator. A cuda-rewritten graph therefore produces stream-less autograd nodes, and AOT autograd's backward trace inside compile_fx trips opt_ready_stream && opt_parent_stream (engine.cpp:1085) -- the cause of 8 of the 11 test failures. Verified by differential test: eager backward on plain cuda fails identically with torch.compile never involved. Registration surface (device_interface.py, inductor_codegen.py): * GPU_TYPES gains "flagos" in place -- is_gpu() is a membership test on that list object, and without it inductor takes the C++/CPU codegen path and never emits Triton. get_gpu_type()'s functools cache is primed while the list is narrowed, since it asserts at most one GPU type is available and the torch.cuda shim reports available too. * DeviceInterface subclass: device state from torch.flagos, hardware properties from torch.cuda (same physical GPU, same allocator). * DeviceProperties.create reports flagos as cuda at the Triton boundary. Triton's NVIDIA backend hard-checks target.backend == "cuda", so a literal "flagos" finds 0 compatible backends. Inductor already does this rewrite in the opposite direction for ROCm (hints.py:149). * Device op overrides + scheduling/wrapper codegen: the stock CUDA/Triton pipeline under the "flagos" key, also published on torch.flagos for inductor's official PrivateUse1 hook. Two generated-kernel bugs that only surface under compilation: * detach re-dispatched into itself. The kernel called at::detach(self), also registered on PrivateUse1, so it dispatched straight back. Eager hid the recursion because DeviceBoxingGuard rewrites self's device metadata first; under FakeTensor it cannot, since the Python dispatch key sits above the backend key. Dynamo traces every nn.Linear through detach, so this was a stack-overflow segfault at trace time. Now emits at::native::detach (NATIVE_DIRECT_VIEW_OPS). * gen_inplace passed only plain at::Tensor args to DeviceBoxingGuard, so clamp_.Tensor handed unboxed flagos min/max to a CUDA self and crashed. Optionals are now materialized into holders, matching gen_functional_pure. Both regressions are covered by tests that were confirmed to fail (segfault at the exact asserting line) against a build with the fixes reverted. CPU-torch wheel accommodations, since torch.cuda's Python layer was frozen without CUDA: re-attach CudaInterface.get_raw_stream (binding exists, the import-time _is_compiled() probe left it None), route torch.cuda.memory_* to the flagos allocator that backs the same pool, hand out flagos Event/Stream in place of the dummy base classes, force triton.cudagraphs off (torch.cuda.CUDAGraph raises on construction) and use_static_cuda_launcher off (not built). flagos_compile_backend now accepts the mode/options/dynamic kwargs dynamo forwards to named backends and expands them into compile_fx config_patches, rather than mutating inductor's global config. Tests: test_compile.py 12 passed / 1 skipped (was 2 passed / 8 failed); test_clamp_dispatch.py 15 passed; ops dispatch sweep 358 passed; test_ops.py 58 passed; allocator/factory/fallback/unit 73 passed. Docs updated to drop the device-aliasing description and the unmeasured performance-parity figures; benchmarking remains open work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zhaoyinglia
approved these changes
Aug 5, 2026
lvyufeng
pushed a commit
to lvyufeng/PyTorch-Plugin-FL
that referenced
this pull request
Aug 6, 2026
… flagos The torch.compile integration merged in flagos-ai#41 only ever compiled single-Linear models in its tests, which need neither autotuning nor more than one Triton kernel. Two independent failures hid behind that. Both reproduce on any graph with a couple of stacked Linears or a LayerNorm. Autotuning needs a constructible event. InductorBenchmarker.get_event_pairs times candidate configs with torch.cuda.Event(enable_timing=True). In the CPU-only wheel this build pairs with an external libtorch_cuda.so, that binding was never compiled, so torch.cuda substitutes a placeholder from torch._utils._dummy_type whose __new__ raises "Tried to instantiate dummy base class Event". flagos.Event subclassed it and inherited the failure. flagos.Event now picks its base class by lineage: on a vendor torch build it still subclasses torch.cuda.Event, and when that is a dummy it subclasses the device-agnostic torch.Event, which dispatches record/block/query/ elapsedTime to c10::flagos::DeviceGuardImpl (csrc/runtime/guard.h). Timing stays a real device measurement, and since every vendor under csrc/runtime/accelerator/ implements that ABI, the fallback is portable rather than NVIDIA-specific. Note the fix has to land here: patching triton.testing.do_bench does not help, because inductor reaches the benchmarker through triton_heuristics.benchmark_all_configs -> bench -> benchmarker.benchmark_gpu, not through do_bench. Compile workers need torch_fl. Inductor's default worker_start_method, "subprocess", starts workers as a bare `sys.executable -m torch._inductor.compile_worker` that imports only torch and triton. flagos lives behind PrivateUse1, so such a worker has no accelerator: triton's CudaDriver.is_active() asks torch.cuda.is_available(), gets False, and the worker dies with "Could not find an active GPU backend". "fork" inherits this process, torch_fl included, so workers come up already seeing the device -- and compilation stays parallel, unlike compile_threads = 1 (Qwen3-0.6B: 31.9s forked vs 40.8s serial). Both overrides are scoped to this build by probing for a missing torch._C CUDA binding, so a vendor torch install keeps inductor's defaults. tests/integration/test_compile_autotune.py guards both: stacked Linears, normalizations, reductions, multi-kernel backward, dynamic shapes and max-autotune, plus a direct check that the autotuner's own Event call works. On a cleared TORCHINDUCTOR_CACHE_DIR, 7 of its 8 tests fail before these changes. CI ran no compile tests at all, so .github/configs/cuda.yml now runs both compile files -- in one pytest invocation on purpose. The worker pool is created lazily and shared, so a file run on its own can be served entirely before the pool spins up, which is exactly how the worker failure stayed hidden until a multi-file run reproduced it. Measured on one A100, fp32, compiled vs eager on flagos: Qwen3-0.6B forward 2.24x (35.6ms -> 15.9ms, numerics matching eager at rtol/atol 2e-2), elementwise chain 4096x4096 9.18x, transformer block 1.41x, matmul-bound MLP 1.06x, and 0.92x at 64x512 where launch overhead exceeds the saving. tests/perf/bench_compile.py had never been run and could not be: it called torch.gelu (nonexistent), read torch.os.environ, recognised only the "privateuseone" spelling of the device, and imported torch before torch_fl -- which the docs now state as a hard requirement, since torch_fl preloads the libtorch_cuda.so that torch.cuda depends on. The test suite gets away without it because conftest imports torch_fl during collection. Also documents a third bug found while benchmarking and left unfixed: convolutions do not compile. Inductor prefers channels_last for conv on GPU, and while the flagos conv kernel honours that layout, its fake/meta kernel still predicts contiguous strides, so inductor rejects the graph on a stride mismatch. Eager never hits it, since it is the layout pass that produces a channels_last input. Reproduce with bench_compile.py --model=conv.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Enables
torch.compileon the flagos device by registering flagos withTorchInductor as a first-class GPU device. The traced graph is handed to
compile_fxunchanged, and inductor emits Triton kernels that operate on flagostensors directly — no conversion to cuda, no copy at the graph boundary.
This works because flagos runs on the physical GPU that
torch.cudadescribes:its allocator delegates to
c10::cuda::CUDACachingAllocator, so a flagostensor's storage already is CUDA memory.
Why not the device-aliasing approach
The first revision of this PR rewrote the graph and its example inputs to cuda
before calling
compile_fx. That is not just a copy per call — it breaksbackward.
at::getAccelerator()is PrivateUse1/flagos in this build, andtorch::autograd::Node::stream()only yields a stream when a node's input devicetype equals the accelerator. A cuda-rewritten graph therefore produces
stream-less autograd nodes, and AOT autograd's backward trace inside
compile_fxtrips
opt_ready_stream && opt_parent_stream(engine.cpp:1085). That was thecause of 8 of the 11 test failures.
Confirmed by differential test: eager backward on plain cuda fails identically,
with
torch.compilenever involved.Registration surface
GPU_TYPES.append("flagos")is_gpu()is a membership test on that list object; without it inductor takes the C++/CPU codegen path and never emits Triton.get_gpu_type()'s cacheregister_interface_for_devicetorch.flagos, hardware properties fromtorch.cuda(same physical GPU).DeviceProperties.createwraptarget.backend == "cuda", so a literal"flagos"finds 0 compatible backends. Inductor already does this in reverse for ROCm (hints.py:149).register_device_op_overridesCUDADeviceOpOverrides— attributes present on the base class never reach__getattr__delegation.register_backend_for_device"flagos"key.Two generated-kernel bugs that only surface under compilation
detachre-dispatched into itself. The generated kernel calledat::detach(self), which is also registered on PrivateUse1, so it dispatchedstraight back — infinite recursion. Eager hid this because
DeviceBoxingGuardrewrites
self's device metadata first; under FakeTensor it cannot, since thePython dispatch key sits above the backend key. Dynamo traces every
nn.Linearthroughdetach, so this was a stack-overflow segfault at tracetime. Now emits
at::native::detach(NATIVE_DIRECT_VIEW_OPS).gen_inplacedidn't boxoptional<Tensor>.clamp_.Tensorhanded unboxedflagos
min/maxto a CUDAselfand crashed. Optionals are now materializedinto holders, matching
gen_functional_pure.Both are covered by tests confirmed to fail (segfault at the exact asserting
line) against a build with the fixes reverted.
CPU-torch wheel accommodations
This build pairs a CPU-only pip torch with an external
libtorch_cuda.so, soseveral
torch.cudaPython bindings are missing:CudaInterface.get_raw_streamre-attached — the binding exists, but theimport-time
_is_compiled()probe left itNone.torch.cuda.memory_*routed to the flagos allocator backing the same pool.Event/Streamin place of the dummy base classes.triton.cudagraphs = False(torch.cuda.CUDAGraphraises on construction) anduse_static_cuda_launcher = False(not built).flagos_compile_backendalso accepts themode/options/dynamickwargs dynamoforwards to named backends, expanding them into
compile_fxconfig_patchesrather than mutating inductor's global config.
Testing
Rebased onto
flagos/main(0700b61) and re-verified end to end on A100 in thetorch-fl-210env:tests/integration/test_compile.pytests/integration/ops/test_clamp_dispatch.pytests/integration/ops/(full sweep)tests/integration/ops/test_rng_dispatch.pytest_ops/allocator/factory/fallback_trace/clone_dispatchtests/unitThe rebase brought in the RNG generator-injection work (#39, #49), which touches
the same
scripts/codegen_ops.pytemplates; the conflict was resolved keepingboth, and
python scripts/codegen_ops.pywas verified to reproduce the committedcuda_kernels.ccbyte-for-byte.Open work
FLAGOS_USE_FLAGTREE=1, off by default)revision of this PR quoted speedup figures that were measured under the old
device-aliasing design, so they no longer describe this code and have been
dropped rather than restated.
🤖 Generated with Claude Code